Papers with moral alignment

2 papers
The Greatest Good Benchmark: Measuring LLMs’ Alignment with Utilitarian Moral Dilemmas (2024.emnlp-main)

Copied to clipboard

Challenge: Our analysis across 15 diverse LLMs reveals consistently encoded moral preferences that diverge from established moral theories and lay population moral standards.
Approach: They propose to evaluate the moral judgments of large language models using utilitarian dilemmas to determine their moral alignment.
Outcome: The findings highlight the ‘artificial moral compass’ of Large Language Models, offering insights into their moral alignment.
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) have advanced multimodal learning, driving progress in cross-modal reasoning.
Approach: They propose to examine moral robustness of vision-language models by analyzing their moral stances under multimodal perturbations.
Outcome: The proposed model-agnostic multimodal perturbations expose VLMs to a variety of moral vulnerabilities, including a sycophancy trade-off where stronger instruction-following models are more susceptible to persuasion.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations